Papers with German language
Neural OCR Post-Hoc Correction of Historical Corpora (2021.tacl-1)
Copied to clipboard
| Challenge: | Optical character recognition (OCR) is crucial for a deeper access to historical collections. |
| Approach: | They propose a neural approach based on a combination of recurrent (RNN) and deep convolutional network (ConvNet) to correct OCR transcription errors. |
| Outcome: | The proposed model reduces the word error rate of 32.3% by more than 89% on a historical book corpus in German language. |
PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text (2021.eacl-demos)
Copied to clipboard
| Challenge: | Prior punctuation restoration methods have focused on using lexical features, prosodic features or combination of both. |
| Approach: | They propose a multitask modeling approach to restore punctuation in multiple high resource languages using acoustic models and a computational model. |
| Outcome: | The proposed system can restore punctuation in Germanic, Romanic and low resource languages without extensive knowledge of grammar or syntax. |
A Corpus for Argumentative Writing Support in German (2020.coling-main)
Copied to clipboard
| Challenge: | In today's world most information is readily available. Consequently, the sole reproduction of information is losing attention. |
| Approach: | They propose an annotation approach to capture claims and premises of arguments and their relations in student-written peer reviews on business models in german language. |
| Outcome: | The proposed annotation scheme guides annotators to moderate agreement with the proposed scheme on 50 persuasive student-written peer reviews on business models. |
Robustness Evaluation of the German Extractive Question Answering Task (2025.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for Question Answering systems only include EM and F1 scores, but they overlook critical factors for the deployment of QA systems. |
| Approach: | They propose to define an evaluation method specifically tailored to the German language to evaluate the robustness of German QA models. |
| Outcome: | The proposed method extends existing methods to German language . it shows that all models are vulnerable to character-level perturbations . |
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)
Copied to clipboard
| Challenge: | lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems. |
| Approach: | They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation. |
| Outcome: | The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers. |
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented. |
| Approach: | They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods. |
| Outcome: | The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents. |
Supporting Land Reuse of Former Open Pit Mining Sites using Text Classification and Active Learning (2021.acl-long)
Copied to clipboard
Christopher Schröder, Kim Bürgl, Yves Annanias, Andreas Niekler, Lydia Müller, Daniel Wiegreffe, Christian Bender, Christoph Mengs, Gerik Scheuermann, Gerhard Heyer
| Challenge: | open pit mines left many regions worldwide inhospitable or uninhabitable . aforementioned information has to be acquired to ensure safety and validity of land reuse . |
| Approach: | They propose a workflow for supporting the post-mining management of former open pit mines in the eastern part of Germany . they use active learning to perform multi-label sentence classification for two categories of restrictions and seven categories of topics . |
| Outcome: | The proposed system supports the post-mining management of former lignite open pit mines in the eastern part of Germany. |
GRhOOT: Ontology of Rhetorical Figures in German (2022.lrec-1)
Copied to clipboard
| Challenge: | GRhOOT is a domain ontology of rhetorical figures in the German language . the goal is to allow for easier detection of non-literal language based tasks . |
| Approach: | GRhOOT is a domain ontology of 110 rhetorical figures in the german language . the goal is to allow for easier detection and sentiment analysis . |
| Outcome: | The ontology of rhetorical figures in the German language is based on 110 rhetorical figure domains . the goal is to make the ontologies more accurate and to allow for easier detection . |
SuperGLEBer: German Language Understanding Evaluation Benchmark (2024.naacl-long)
Copied to clipboard
| Challenge: | a new set of German-pretrained models are being released, but no established, diverse and systematic evaluation suite is available for them. |
| Approach: | They assemble a Natural Language Understanding benchmark suite for the German language and evaluate 10 existing German-pretrained models. |
| Outcome: | The proposed benchmark suite evaluates 10 German-pretrained models on 29 tasks . the results show that encoder models are good choices for most tasks, but not all . |
A Joint Approach to Compound Splitting and Idiomatic Compound Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | Compounding is a common word-formation process in Germanic languages . high productivity and low corpus frequency of compounds increase vocabulary size . |
| Approach: | They develop a deep learning-based approach to noun compound splitting and idiomatic compound detection for the German language. |
| Outcome: | The proposed approach outperforms the current state of the art in noun compound splitting and idiomatic compound detection for the German language. |
Modeling Persuasive Discourse to Adaptively Support Students’ Argumentative Writing (2022.acl-long)
Copied to clipboard
| Challenge: | Argumentation is an omnipresent rudiment of daily communication and thinking . humans struggle to develop argumentation skills due to a lack of individual and instant feedback in their learning process. |
| Approach: | They propose an argumentation annotation approach to model argumentative discourse in student-written business model pitches and embed it into an adaptive writing support system for students that provides individual argumentation feedback. |
| Outcome: | The proposed method annotates a corpus of 200 business model pitches in german and measures their self-efficacy and ease-of-use in a real-world writing exercise. |
Abstract Text Summarization: A Low Resource Challenge (D19-1)
Copied to clipboard
| Challenge: | Existing datasets for multilingual text summarization are difficult to construct and lack of human knowledge and language processing abilities in computers makes text summaries a challenging task. |
| Approach: | They propose an iterative data augmentation approach which uses synthetic data along with the real summarization data for the German language. |
| Outcome: | The proposed system improves on the development and test sets on the German language text using the state-of-the-art “Transformer” model. |
FinCorpus-DE10k: A Corpus for the German Financial Domain (2024.lrec-main)
Copied to clipboard
| Challenge: | a predominantly German corpus of financial documents is available for the first time . financial text is characterized by a unique vocabulary with implications including sentiment analysis . |
| Approach: | They propose a predominantly German financial corpus comprising 12.5k PDF documents . they hope it will fill this gap and foster further research in the financial domain . |
| Outcome: | The proposed corpus is the first non-email German financial corpus available . it aims to provide insights into financial discourse in the German language and multilingually. |
German SRL: Corpus Construction and Model Training (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing semantic role annotation resources are lacking for German. |
| Approach: | They propose a translation-based approach to train German semantic role models using semantic annotations and alignment models. |
| Outcome: | The proposed method achieves competitive evaluation scores, but avoids limitations of previous approaches. |
From Witch’s Shot to Music Making Bones - Resources for Medical Laymen to Technical Language and Vice Versa (2020.lrec-1)
Copied to clipboard
| Challenge: | Information we share online unveils directly or indirectly information about our lifestyle and health situation. |
| Approach: | They propose a dataset which annotates medical laymen and technical expressions in a patient forum and a set of medical synonyms and definitions. |
| Outcome: | The proposed dataset annotates medical laymen and technical expressions in a patient forum along with a set of medical synonyms and definitions. |
Summarization Corpora of Wikipedia Articles (2020.lrec-1)
Copied to clipboard
| Challenge: | Using Wikipedia articles, we extract summarization data for other languages. |
| Approach: | They propose a process to extract Wikipedia summarization corpora and apply it to the German language. |
| Outcome: | The proposed method can be applied to the German language and compares to baselines. |
PopAut: An Annotated Corpus for Populism Detection in Austrian News Comments (2024.lrec-main)
Copied to clipboard
| Challenge: | Populism is a phenomenon that is noticeably present in political landscapes worldwide . prior work on populism analysis focused on analyzing populist content expressed by politicians . |
| Approach: | They present a corpus of news comments annotated for populism in the german language . they use machine learning to detect populist comments in text . |
| Outcome: | The proposed corpus outperforms existing dictionaries for populism detection in text . it features 1,200 comments collected between 2019-2021 . |
Using Pre-Trained Language Models in an End-to-End Pipeline for Antithesis Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Rhetorical figures are a "departure from the normal usage" of language . features of metaphors, irony and sarcasm enhance performance of several NLP tasks. |
| Approach: | They propose a pipeline approach to detect rhetorical figures using large language models by splitting text into phrases and identifying parallel phrases with a syntactically parallel structure. |
| Outcome: | The proposed approach outperforms state-of-the-art methods by an F1 score of 65.11 %. |